CLI Tools
Health data engineering happens in the terminal more often than the slides suggest. These are the command-line tools that repay learning.
FHIR
HL7 FHIR Validator — the reference validator, run against resources, profiles and implementation guides. The fastest way to settle an argument about whether a resource is conformant.
java -jar validator_cli.jar patient.json -version 4.0.1 -ig hl7.fhir.us.core
SUSHI — compiles FHIR Shorthand (FSH) into profiles, extensions and value sets. Writing profiles as text you can diff and review beats editing JSON by hand.
sushi build .
IG Publisher — builds a full implementation guide site from FSH output and narrative pages.
fhir-py / fhirclient / fhir.resources — Python libraries with CLI-friendly entry points for scripted resource work.
HTTP and APIs
curl — the baseline. If it does not work in curl, it does not work.
curl -s -H "Accept: application/fhir+json" \
"https://example.org/fhir/metadata" | jq '.rest[0].resource[].type'
httpie — friendlier syntax for exploratory work.
jq — the essential JSON processor: filter, reshape, extract.
jq '.entry[].resource | {id, birthDate}' bundle.json
yq — the same idea for YAML.
See APIs for monitoring and observability tooling.
Data wrangling
csvkit — csvlook, csvstat, csvsql for inspecting and querying CSV
extracts without opening a spreadsheet.
Miller (mlr) — CSV/TSV/JSON processing with named fields; excellent for
reshaping export files.
DuckDB — SQL over CSV and Parquet files directly from the shell, fast enough for most analysis of routine health data extracts.
duckdb -c "SELECT district, count(*) FROM 'facilities.csv' GROUP BY 1"
pandas via a short script — when the transformation stops fitting on one line.
Databases and deployment
- psql — PostgreSQL client;
\copy,\timingandEXPLAIN ANALYZEare the workhorses of query tuning. - pg_dump / pg_restore — backups and environment refreshes.
- docker and docker compose — see Docker.
- rclone, rsync — moving extracts and backups between environments.
Working safely
- Never pass credentials as command-line arguments. They land in shell history and process listings; use environment variables or credential files with restricted permissions.
- De-identify before you experiment. Pull synthetic or de-identified data for exploratory work — see health data.
- Script it once you have run it twice. Repeated manual pipelines drift.
- Keep scripts in version control, including the ad-hoc ones. Today's quick fix is next quarter's undocumented dependency.
Related
- Docker — running these tools reproducibly
- Analytics tools
- FHIR servers